Search for: All records

Creators/Authors contains: "Peng, J."

« Prev Next »

Total Resources

99

Resource Type
Conference Paper

7

Conference Proceeding

0

Dataset

0

Journal Article

92

Workshop Report

0

Availability
Full Text / Resource Available

97

Citation Only

2

Save Results
Excel (limit 2000)
CSV (limit 5000)
XML (limit 5000)

Have feedback or suggestions for a way to improve these results?
!

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

CAT s are Fuzzy PETs : A Corpus and Analysis of Potentially Euphemistic Terms

Gavidia, M. ; Lee, P. ; Feldman, A. ; Peng, J. ( January 2022 , arXiv preprint arXiv:2205.02728.)

Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nevertheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, which may be useful for future applications. We also discuss the results of multiple analyses run on the corpus. Firstly, we find that sentiment analysis on the euphemistic texts supports that PETs generally decrease negative and offensive sentiment. Secondly, we observe cases of disagreement in an annotation task, where humans are asked to label PETs as euphemistic or not in a subset of our corpus text examples. We attribute the disagreement to a variety of potential reasons, including if the PET was a commonly accepted term (CAT).
more » « less
Full Text Available
Is self-supervised learning more robust than supervised learning?

Zhong, Y. ; Tang, H. ; Chen, J. ; Peng, J. ; Wang, Y.-X. ( January 2022 , Proc ICML Workshop on Pre-training)

Full Text Available
Measurement of flavor asymmetry of the light-quark sea in the proton with Drell-Yan dimuon production in p+p and p+d collisions at 120 GeV

https://doi.org/10.1103/PhysRevC.108.035202

Dove, J. ; Kerns, B. ; Leung, C. ; McClellan, R. E. ; Miyasaka, S. ; Morton, D. H. ; Nagai, K. ; Prasad, S. ; Sanftl, F. ; Scott, M. B. ; et al ( September 2023 , Physical Review C)

Free, publicly-accessible full text available September 1, 2024
Distributed code for semantic relations predicts neural similarity during analogical reasoning

Chiang, J. N. ; Peng, J. ; Lu, H. ; Holyoak, K. J. ; Monti, M. M. ( March 2021 , Journal of cognitive neuroscience)
null (Ed.)
The ability to generate and process semantic relations is central to many aspects of human cognition. Theorists have long debated whether such relations are coarsely coded as links in a semantic network or finely coded as distributed patterns over some core set of abstract relations. The form and content of the conceptual and neural representations of semantic relations are yet to be empirically established. Using sequential presentation of verbal analogies, we compared neural activities in making analogy judgments with predictions derived from alternative computational models of relational dissimilarity to adjudicate among rival accounts of how semantic relations are coded and compared in the brain. We found that a frontoparietal network encodes the three relation types included in the design. A computational model based on semantic relations coded as distributed representations over a pool of abstract relations predicted neural activities for individual relations within the left superior parietal cortex and for second-order comparisons of relations within a broader left-lateralized network.
more » « less
Full Text Available
You Don’t Say....Linguistic Features in Sarcasm Detection

Ducret M. ; Kruse L. ; Martinez C. ; Feldman A. ; Peng, J. ( January 2021 , CLIC-IT 2021: Seventh Italian Conference on Computational Linguistics Bologna)

We explore linguistic features that contribute to sarcasm detection. The linguistic features that we investigate are a combination of text and word complexity, stylistic and psychological features. We experiment with sarcastic tweets with and without context. The results of our experiments indicate that contextual information is crucial for sarcasm prediction. One important observation is that sarcastic tweets are typically incongruent with their context in terms of sentiment or emotional load.
more » « less
Full Text Available
Pixel contrastive-consistent semi-supervised semantic segmentation

https://doi.org/10.1109/ICCV48922.2021.00718

Zhong, Y. ; Yuan, B. ; Wu, H. ; Yuan, Z. ; Peng, J. ; Wang, Y.-X. ( January 2021 , International Conference on Computer Vision)

Full Text Available
Linguistic Fingerprints of Internet Censorship: the Case of Sina Weibo

Ng Kei Y. ; Feldman A. ; Peng, J ( January 2020 , Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-20))

This paper studies how the linguistic components of blogposts collected from Sina Weibo, a Chinese microblogging platform, might affect the blogposts’ likelihood of being censored. Our results go along with King et al. (2013)’s Collective Action Potential (CAP) theory, which states that a blogpost’s potential of causing riot or assembly in real life is the key determinant of it getting censored. Although there is not a definitive measure of this construct, the linguistic features that we identify as discriminatory go along with the CAP theory. We build a classifier that significantly outperforms non-expert humans in predicting whether a blogpost will be censored. The crowdsourcing results suggest that while humans tend to see censored blogposts as more controversial and more likely to trigger action in real life than the uncensored counterparts, they in general cannot make a better guess than our model when it comes to ‘reading the mind’ of the censors in deciding whether a blogpost should be censored. We do not claim that censorship is only determined by the linguistic features. There are many other factors contributing to censorship decisions. The focus of the present paper is on the linguistic form of blogposts. Our work suggests that it is possible to use linguistic properties of social media posts to automatically predict if they are going to be censored.
more » « less
Full Text Available
Neural Network Prediction of Censorable Language

Ng Kei Y ; Feldman A ; Peng J., and C. ( January 2019 , Proceedings of the 3rd Workshop on NLP and Computational Social Science (NLP+CSS) held in conjunction with NAACL 2019)

Internet censorship imposes restrictions on what information can be publicized or viewed on the Internet. According to Freedom House’s annual Freedom on the Net report, more than half the world’s Internet users now live in a place where the Internet is censored or restricted. China has built the world’s most extensive and sophisticated online censorship system. In this paper, we describe a new corpus of censored and uncensored social media tweets from a Chinese microblogging website, Sina Weibo, collected by tracking posts that mention ‘sensitive’ topics or authored by ‘sensitive’ users. We use this corpus to build a neural network classifier to predict censorship. Our model performs with a 88.50% accuracy using only linguistic features. We discuss these features in detail and hypothesize that they could potentially be used for censorship circumvention.
more » « less
Full Text Available
Publisher Correction: The asymmetry of antimatter in the proton

https://doi.org/10.1038/s41586-022-04707-z

Dove, J. ; Kerns, B. ; McClellan, R. E. ; Miyasaka, S. ; Morton, D. H. ; Nagai, K. ; Prasad, S. ; Sanftl, F. ; Scott, M. B. ; Tadepalli, A. S. ; et al ( April 2022 , Nature)

Full Text Available
Linguistic Characteristics of Censorable Language on Sina Weibo

Ng Kei Y. ; Feldman A ; Peng J. and Leberknight, C. ( January 2018 , Proceedings of The COLING 1st Natural Language Processing for Information Freedom workshop)

This paper investigates censorship from a linguistic perspective. We collect a corpus of censored and uncensored posts on a number of topics, build a classifier that predicts censorship decisions independent of discussion topics. Our investigation reveals that the strongest linguistic indicator of censored content of our corpus is its readability.
more » « less
Full Text Available

« Prev Next »